Papers with text processing

18 papers
An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks (2020.aacl-main)

Copied to clipboard

Challenge: Traditionally, tokenization is the very first step in most text processing works.
Approach: They propose to use morphological segmentation followed by BPE for Korean NLP tasks . they empirically examine what is the best tokenization strategy for Korean to/from English .
Outcome: The proposed approach is best for Korean to/from English machine translation and natural language understanding tasks.
Octopus: On-device language model for function calling of software APIs (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are pivotal for advanced text processing and generation.
Approach: They propose a framework to train on-device Large Language Models optimized for invoking software APIs.
Outcome: The proposed model outperforms GPT-4 in API calling tasks while maintaining inference speed.
On-Device Neural Language Model Based Word Prediction (C18-2)

Copied to clipboard

Challenge: Currently, on-device keyboards have limited memory and response time for word prediction . a proposed on-device neural language model based word prediction method is available for mobile devices .
Approach: They propose an on-device neural language model based word prediction method that optimizes run-time memory and provides a real-time prediction environment.
Outcome: The proposed model outperforms existing methods for word prediction in keystroke savings and word prediction rate and has been commercialized.
VISPool: Enhancing Transformer Encoders with Vector Visibility Graph Neural Networks (2024.findings-acl)

Copied to clipboard

Challenge: Existing graph-based graph construction methods rely on static graphs and are not scalable with increasing document and word counts.
Approach: They propose a dynamic graph construction method based on vector visibility graphs (VVGs) they propose scalable model architecture that integrates VVG convolutional networks into transformer pipelines.
Outcome: The proposed model outperforms baseline models on the GLUE benchmark datasets.
Finding the Law: Enhancing Statutory Article Retrieval via Graph Neural Networks (2023.eacl-main)

Copied to clipboard

Challenge: Statutory article retrieval (SAR) is a promising application of legal text processing.
Approach: They propose a graph-augmented dense statute retriever model that incorporates the structure of legislation via a neural network to improve density retrieval performance.
Outcome: The proposed model outperforms baselines on a real-world expert-annotated dataset.
Learning Context-Sensitive Convolutional Filters for Text Processing (D18-1)

Copied to clipboard

Challenge: Convolutional neural networks (CNNs) are a popular building block for natural language processing . despite their success, most existing CNN models share the same learned set of filters for all input sentences.
Approach: They propose to use a meta network to learn context-sensitive convolutional filters for text processing by using a bidirectional filter generation mechanism.
Outcome: The proposed framework outperforms standard and attention-based CNN models on four different tasks.
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects (2026.acl-long)

Copied to clipboard

Challenge: Research on cross-dialectal transfer from a standard to a non-standard dialect variety has typically focused on text data.
Approach: They compare standard-to-dialect transfer in three settings: text models, speech models, and cascaded systems where speech first gets automatically transcribed and then further processed by a text model.
Outcome: The proposed model performs best on German dialect data while the text-only model perform best on the standard data.
A unified approach to sentence segmentation of punctuated text in many languages (2021.acl-long)

Copied to clipboard

Challenge: Existing tools for segmenting punctuated text in many languages are limited in their language coverage and evaluation is ad hoc.
Approach: They propose a new context-based modeling approach that can be trained on noisily-annotated data.
Outcome: The proposed model exceeds baselines set by existing methods on English corpora and performs well on average on new multilingual evaluation set.
Neural Topic Model with Reinforcement Learning (D19-1)

Copied to clipboard

Challenge: Experimental results show superior performance on perplexity and topic coherence measures compared to state-of-the-art topic models.
Approach: They propose to incorporate topic coherence measures as reward signals to guide the learning of a VAE-based topic model.
Outcome: The proposed model is able to separating background words dynamically from topic words eliminating the pre-processing step of filtering infrequent and/or top frequent words, typically required for learning traditional topic models.
Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent years have witnessed significant advancements in large language models (LLMs) but still struggle with integrating vision and audio.
Approach: They propose a self-knowledge distillation method to improve vision-audio capabilities of OLLMs by learning from the vision-text components.
Outcome: The proposed method improves vision-audio capabilities of OLLMs by learning from vision-text components, which improves interaction between audio and images and results in improved performance on multimodal tasks.
HIT - A Hierarchically Fused Deep Attention Network for Robust Code-mixed Language Representation (2021.findings-acl)

Copied to clipboard

Challenge: linguistics and morphology of resource-short code-mixed texts remain a key challenge in text processing.
Approach: They propose a hierarchical transformer-based framework that captures the semantic relationship among words and hierarchically learns sentencelevel semantics using a fused attention mechanism.
Outcome: The proposed framework improves on one European and five Indic languages on four NLP tasks on eleven datasets.
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are generalist agents capable of operating within complex environments.
Approach: They propose a class of tools that can serve as a middleware layer shielding LLMs from environmental complexity.
Outcome: The proposed tool can shield the LLM from environmental complexity in two representative complex environments.
Spherical Latent Spaces for Stable Variational Autoencoders (D18-1)

Copied to clipboard

Challenge: Variational autoencoders use a multivariate Gaussian latent variable to capture latent structure in data.
Approach: They propose a variational autoencoder which uses a latent distribution instead of Gaussian . they find that the variational posterior averts the KL collapse by a fixed hyperparameter .
Outcome: The von Mises-Fisher distribution averts the KL collapse and gives better likelihoods than Gaussian models across a range of modeling conditions.
Multilingual Culture-Independent Word Analogy Datasets (2020.lrec-1)

Copied to clipboard

Challenge: In text processing, deep neural networks use word embeddings as an input.
Approach: They propose to use benchmark datasets to compare the quality of word embeddings in text processing . they use a word analogy task in Croatian, English, Estonian, Finnish, Latvian, Lithuanian, Russian, Slovenian, and Swedish .
Outcome: The proposed datasets are culturally independent and cross-lingual for the languages used.
Sequential Learning of Convolutional Features for Effective Text Classification (D19-1)

Copied to clipboard

Challenge: Existing models for text classification have largely ignored convolution filters and max pooling . text classification is one of the major applications of natural language processing .
Approach: They propose a convolutional attentive recurrent network model which uses convolution filters and max pooling to improve text classification.
Outcome: The proposed model outperforms existing convolutional models on text classification tasks.
Consonant is all you need: a compact representation of English text for efficient NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: In natural language processing, the representation of text plays a crucial role in various tasks such as language modeling, sentiment analysis, and machine translation.
Approach: They propose a method to represent English text with only consonants that is more discriminative than vowels and a technique to retrieve vowel information from it.
Outcome: The proposed representation significantly reduces the overall memory and compute footprint required for storing and processing textual data.
C²RBench: A Chinese Complex Reasoning Benchmark for Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks often fail to capture complex multi-step reasoning demands inherent in real-world scenarios.
Approach: They propose a benchmark to evaluate multi-step, multimodal advanced reasoning of large language models.
Outcome: The proposed benchmark exceeds existing benchmarks in cognitive complexity and accuracy by over 90% . it features 1,115 carefully curated Chinese tasks organized into eight domain-specific subsets . evaluations of 20 LLMs and 24 multimodal large language models reveal critical performance gaps .
Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks assess basic knowledge breadth or lexical understanding, failing to capture higher-order skills that are central to historical research.
Approach: They propose a benchmark anchored in the Chinese Imperial Examination system that assesses historical knowledge and lexical understanding.
Outcome: The new benchmark aims to assess the ability of LLMs to process historical materials and documents.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations